Sarvam AI2026-09-26 22:04:57Sarvam AI launches Saaras V4 speech-to-text model for 22 Indian languagesIndian AI company Sarvam AI has introduced Saaras V4, a new speech recognition model that supports all 22 official Indian languages as well as global English accents. The model is available through an API and includes real-time streaming transcription, batch processing, and speaker diarization. Sarvam AI lists real-time transcription pricing at 30 Indian rupees per hour. Saaras V4 uses an encoder-decoder architecture, with the decoder built on Sarvam-3B, the company’s in-house 3 billion-parameter hybrid state space language model. It offers five output modes: transcription, verbatim records, code-mixed output, romanized transliteration, and translation. The company also added a keyword prompting feature designed to improve recognition accuracy for specific terms. According to Sarvam AI, Saaras V4 posted the best results across multiple benchmarks and recorded a lower error rate than Deepgram Nova-3 and GPT-4o Transcribe on noisy audio datasets. Those figures, however, were reported by the company itself, and no independent replication results have been published so far. The report was cited by MarkTechPost.20